Skip to content

CI: owner-finals (yas-run/yas#93) - #46

Closed
pcarrier wants to merge 4 commits into
ci-basefrom
owner-finals
Closed

pcarrier wants to merge 4 commits into
ci-basefrom
owner-finals

Conversation

@pcarrier

Copy link
Copy Markdown

Fork CI for yas-run#93 (stacked on yas-run#92), on 038bbfb. Not for merging.

A WAIT on a process its session holds no route to looks at it with a temporary
WATCH. That look is refused as CONFLICT while the process's exit is on its way to
its watchers (the record is final only once they all have it), or while the
session's own previous look, such as the one a KILL's CONTROL just made, is still
leaving. So a WAIT right after a KILL of a detached process could answer CONFLICT
instead of the exit: client_host's
a_process_whose_attachment_went_is_waited_for_and_held_by_nothing failed so on a
loaded machine. Such a WAIT now waits for the process catalogue's next change,
with which either settles, and looks again, within its timeout.

a_wait_right_after_a_kill_of_a_detached_process_gets_its_exit (64 rounds of
spawn detachable, detach, KILL, WAIT) failed 1 run in 10 without the change, and
none in 20 with it.

a_watcher_whose_queue_fills_after_the_exit_fails_alone took only output in its
frame loop and panicked on the stdin's progress event, which a slow machine
delivers among the frames; it lets it pass now, and waits for the child's reaping
(Server::wait_reaped, tests only) rather than 500 ms.
A WAIT's temporary look can be refused as CONFLICT because this session's own
look at the process is still bound: a concurrent CONTROL's, or a route that an
eviction failed and has yet to detach (dispatch_outbound fails the route, then
sends the Detach). Such a WAIT waited for the catalogue's next change, which a
Detach doesn't make, so it slept until the process ended and then found nothing
of an ordinary process: NOT_FOUND. It now waits with Manager::wait_native_look,
which also wakes when the record changes and returns once nothing refuses a look:
no exit in flight and no binding of this endpoint.

a_wait_that_finds_its_own_look_still_bound_answers_once_it_goes fails a watcher's
route, WAITs, and detaches the look 200 ms later. Without the change the WAIT
answered NOT_FOUND once the child's second had passed; with it, the exit.

The Exit arm's comment still said such a WAIT answers CONFLICT; it says what it
does now.
An ordinary process belongs to its spawning session. yas-client WAITs for its
exit when the attachment that would report it goes first: a stream reset or
dropped, a DETACH, or a connection that drops and comes back just as the process
exits. That WAIT answered NOT_FOUND whenever the exit came in between. With no
binding of the owner's to take the exit, the record was released and nothing was
kept, since only a detachable process's final is retained. The same happened in
two other cases. When the owner's route failed as the exit was queued, the
adapter dropped the exit because no route was there to take it. When the exit
found the owner's event queue full, it was dropped, and the attachment waited
for it forever.

The owner now finds the exit in each case:
- The adapter records the replay of every exit it is handed, whether its route
  is there or not. NativeEvent::Exit now carries the process handle for that.
- For the owner's endpoint alone, the native side keeps the final of an
  ordinary process whose exit that endpoint missed: it had no binding when the
  exit was queued, or its exit event was dropped before dispatch. WriterGuard
  now tells its action whether its event was dispatched. The endpoint keeps the
  newest exit_replays() of these finals, as many as the adapter keeps replays,
  until it shuts down. WATCH, and so WAIT and ATTACH, finds them after the
  public finals. They are stored before the release moves the catalogue, which
  is what a WAIT that found the exit on its way waits for.
- A binding whose exit can't be queued is evicted, as a binding that falls
  behind is. Its attachment fails and its client WAITs, instead of waiting for
  an exit that won't come.

New tests, each of which failed with NotFound before this change:
- an_owner_whose_look_went_before_the_exit_still_waits_for_it
- an_owner_whose_route_failed_as_the_exit_came_still_waits_for_it
- an_owner_whose_queue_was_full_as_the_exit_came_still_finds_it

an_owner_that_drops_its_stream_after_the_exit_still_waits_for_it no longer
depends on whether its detach comes before or after the exit.

a_watcher_whose_queue_fills_after_the_exit_fails_alone assumed that 80 lines
make 80 frames, which breaks when a loaded reader takes several lines in one
read. Its child now writes each line only after the test has received the last
one, as a FIFO tells it, and runs without stdin, so there are no stdin events.
The test checks each frame, and checks that only the watcher's attachment
failed, evicted.
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Coverage

Crate Lines Functions Regions
alacritty-driver 85.6% (1081/1263) 89.4% (84/94) 89.3% (1746/1955)
browser 25.5% (309/1212) 30.4% (34/112) 27.7% (605/2183)
cli 31.8% (6517/20468) 32.8% (598/1824) 32.9% (9634/29263)
client 66.2% (4898/7399) 68.1% (624/916) 64.5% (6343/9832)
composite-transport 96.3% (526/546) 98.4% (60/61) 96.2% (884/919)
compositor 54.9% (11395/20763) 66.0% (832/1261) 54.7% (15779/28844)
desktop 78.4% (4460/5691) 71.6% (393/549) 75.1% (6211/8267)
edge 62.5% (793/1268) 50.3% (81/161) 57.1% (1027/1799)
fonts 77.3% (1257/1626) 82.7% (129/156) 78.9% (2424/3071)
fssync 85.6% (1602/1872) 85.4% (181/212) 86.7% (2822/3255)
git 70.4% (4462/6334) 68.5% (337/492) 67.3% (6313/9386)
guest 67.3% (6829/10148) 67.0% (488/728) 66.6% (8935/13416)
lsp 78.6% (3571/4542) 80.8% (336/416) 76.6% (5405/7054)
proxy 63.3% (1899/3000) 56.6% (184/325) 63.5% (2929/4613)
runtime-dir 93.6% (117/125) 100.0% (14/14) 94.8% (218/230)
sd-notify 73.9% (68/92) 100.0% (6/6) 83.2% (109/131)
server 70.8% (84296/119092) 73.3% (6126/8352) 68.8% (112948/164204)
ssh 67.2% (708/1054) 75.9% (85/112) 67.2% (1105/1645)
terminal-model 49.9% (314/629) 62.3% (38/61) 50.3% (505/1004)
uplink 94.2% (582/618) 92.6% (50/54) 93.1% (1062/1141)
webrtc-forwarder 33.6% (1321/3932) 44.9% (146/325) 36.0% (2279/6336)
webserver 80.7% (1490/1846) 81.1% (193/238) 83.2% (2551/3067)
website 35.1% (355/1012) 34.4% (53/154) 35.8% (607/1694)
xtask 0.0% (0/5149) 0.0% (0/131) 0.0% (0/9157)
yas 88.8% (27842/31363) 93.8% (2132/2272) 83.9% (43722/52130)
Total 66.4% (166692/251044) 69.4% (13204/19026) 64.8% (236163/364596)

The failed try_send drops the exit's envelope inside send_native, whose guard
keeps the final and releases the record before try_queue_terminal evicts the
binding: the test could see the process gone and the final kept with no
eviction pushed yet. It waits for the eviction's notice (a permit is kept when
it came first).

Under load, it failed 7 of 24 parallel runs of yas_process::tests::a, and 3 of
40 alone, in review; with the wait, none of 24 and none of 40.
@pcarrier

Copy link
Copy Markdown
Author

yas-run#93 merged as dd33f04.

@pcarrier pcarrier closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant